Papers by Munmun De Choudhury
Responsible Evaluation of AI for Mental Health (2026.acl-long)
Copied to clipboard
Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych
| Challenge: | Existing approaches to evaluating AI tools in this domain remain fragmented and inconsistent. |
| Approach: | They propose a taxonomy of AI mental health support types that integrates clinical soundness, social context, and equity to provide a structured basis for evaluation. |
| Outcome: | The proposed framework integrates clinical soundness, social context, and equity, providing a structured basis for evaluation. |
MythTriage: Scalable Detection of Opioid Use Disorder Myths on a Video-Sharing Platform (2025.emnlp-main)
Copied to clipboard
| Challenge: | 108K drug overdose deaths in 2022, according to NIDA . |
| Approach: | They propose a large-scale study of OUD-related myths on YouTube with clinical experts to validate 8 pervasive myths and release an expert-labeled video dataset. |
| Outcome: | The proposed model reduces annotation time and cost by over 76% compared to experts and full LLM labeling. |
Lived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use (2025.naacl-long)
Copied to clipboard
Mohit Chandra, Siddharth Sriraman, Gaurav Verma, Harneet Singh Khanuja, Jose Suarez Campayo, Zihang Li, Michael L. Birnbaum, Munmun De Choudhury
| Challenge: | Adverse Drug Reactions (ADRs) from psychiatric medications are the leading cause of hospitalizations among mental health patients. |
| Approach: | They propose a benchmark and a framework to evaluate LLMs' ability to detect ADRs . they find that LLM responses are more complex and harder to read than experts . |
| Outcome: | The proposed framework evaluates LLMs' ability to detect and deliver expert-aligned mitigation strategies. |
Who’s Asking? Simulating Role-Based Questions for Conversational AI Evaluation (2026.findings-acl)
Copied to clipboard
| Challenge: | Language model users embed personal and social context in their questions. |
| Approach: | They propose a framework for simulating role-based questions using a taxonomy of asker roles for patients, caregivers, practitioners. |
| Outcome: | The proposed framework simulates 15,321 questions that embed each asker role’s goals, behaviors, and experiences. |
Do Large Language Models Align with Core Mental Health Counseling Competencies? (2025.findings-naacl)
Copied to clipboard
Viet Cuong Nguyen, Mohammad Taher, Dongwan Hong, Vinicius Konkolics Possobom, Vibha Thirunellayi Gopalakrishnan, Ekta Raj, Zihang Li, Heather J. Soled, Michael L. Birnbaum, Srijan Kumar, Munmun De Choudhury
| Challenge: | Large language models are promising for mental health, but their alignment with core counseling competencies remains underexplored. |
| Approach: | They propose a benchmark to evaluate 22 general-purpose and medical-finetuned LLMs across five key competencies. |
| Outcome: | The proposed model outperforms generalist models in Intake, Assessment & Diagnosis but struggles with core counseling attributes and professional practice & ethics. |
What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs’ Self-consistency in Closed Domains Via Adversarial Nudge (2026.acl-long)
Copied to clipboard
| Challenge: | Claude exhibits strong resilience, while GPT and Grok demonstrate moderate resilience . open models fall short significantly, while proprietary models exhibit weak resilience compared to open models . |
| Approach: | They propose a framework for stress testing factual fidelity in large language models in the presence of adversarial nudges. |
| Outcome: | The proposed model is robust to adversarial nudges in two closed domains. |
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation frameworks focus on diagnostic accuracy and win-rates and often overlook alignment with patient-specific goals, values, and personalities required for meaningful conversations. |
| Approach: | They propose a framework for synthetically generating realistic, multi-turn mental health sensemaking conversations and a dataset to examine their models in healthcare settings. |
| Outcome: | The proposed framework synthesizes a dataset comprising over 2,200 patient–LLM conversations and evaluates them using human-centric criteria. |
Latent Hatred: A Benchmark for Understanding Implicit Hate Speech (2021.emnlp-main)
Copied to clipboard
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, Diyi Yang
| Challenge: | Existing studies on explicit or overt hate speech have failed to address a more pervasive form based on coded or indirect language. |
| Approach: | They propose a theoretically-justified taxonomy of implicit hate speech and a benchmark corpus with fine-grained labels for each message and its implication. |
| Outcome: | The proposed dataset will serve as a useful benchmark for understanding this multifaceted issue. |
Auditing LLM Responses to Harmful Stereotypes Targeting Mental Health Groups (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored. |
| Approach: | They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities. |
| Outcome: | The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions. |